Everything about Corpus Linguistics totally explained
Corpus linguistics is the
study of language as expressed in
samples
(corpora) or "real world" text. This method represents a
digestive approach to deriving a set of abstract rules by which a
natural language is governed or else relates to another language. Originally done by hand, corpora are largely derived by an automated process, which is corrected.
Computational methods had once been viewed as a
holy grail of
linguistic research, which would ultimately manifest a
ruleset for
natural language processing and
machine translation at a high level. Such hasn't been the case, and since the
cognitive revolution, cognitive linguistics has been largely critical of many claimed practical uses for corpora. However, as
computation capacity and speed have increased, the use of corpora to study language and term relationships en masse has gained some respectability.
The corpus approach runs counter to
Noam Chomsky's view that real language is riddled with performance-related errors, thus requiring careful analysis of small speech samples obtained in a highly controlled laboratory setting.
Corpus linguistics does away with Chomsky's
competence/performance split; adherents believe that reliable language analysis best occurs on field-collected samples, in natural contexts and with minimal experimental interference.
History
A landmark in modern corpus linguistics was the publication by
Henry Kucera and
Nelson Francis of
Computational Analysis of Present-Day American English in 1967, a work based on the analysis of the
Brown Corpus, a carefully compiled selection of current American English, totalling about a million words drawn from a wide variety of sources. Kucera and Francis subjected it to a variety of computational analyses, from which they compiled a rich and variegated opus, combining elements of linguistics, language teaching,
psychology,
statistics, and
sociology. A further key publication was
Randolph Quirk's 'Towards a description of English Usage' (1960, Transactions of the Philological Society, 40-61) in which he introduced
The Survey of English Usage.
Shortly thereafter, Boston publisher
Houghton-Mifflin approached Kucera to supply a million word, three-line citation base for its new
American Heritage Dictionary, the first
dictionary to be compiled using corpus linguistics. The AHD made the innovative step of combining prescriptive elements (how language
should be used) with descriptive information (how it actually
is used).
Other publishers followed suit. The British publisher Collins'
COBUILD monolingual learner's dictionary, designed for users learning
English as a foreign language, was compiled using the
Bank of English.
The
Brown Corpus has also spawned a number of similarly structured corpora: the
LOB Corpus (1960s
British English), Kolhapur (
Indian English), Wellington (
New Zealand English), Australian Corpus of English (
Australian English), the Frown Corpus (
early 1990s American English), and the FLOB Corpus (1990s British English). Other corpora represent many languages, varieties and modes, and include the
International Corpus of English, and the
British National Corpus, a 100 million word collection of a range of spoken and written texts, created in the 1990s by a consortium of publishers, universities (
Oxford and
Lancaster) and the
British Library. For contemporary American English, work has stalled on the
American National Corpus, but the 360 million word
Corpus of American English (1990-present) is now available.
Methods
This means dealing with real input data, where descriptions based on a linguist's intuition are not usually helpful.
Further Information
Get more info on 'Corpus Linguistics'.
|
External Link Exchanges
Do you know how hard it is to get a link from a large encyclopaedia? Well we're different and will prove it. To get a link from us just add the following HTML to your site on a relevant page:
<a href="http://corpus_linguistics.totallyexplained.com">Corpus linguistics Totally Explained</a>
Then simply click through this link from your web page. Our crawlers will verify your link, extract the title of your web page and instantly add a link back to it. If you like you can remove the words Totally Explained and embed the link in article text.
As long as your link remains in place, we'll keep our link to you right here. Please play fair - our crawlers are watching. Your site must be closely related to this one's topic. Any kind of spamming, dubious practises or removing the link will result in your link from us being dropped and, potentially, your whole site being banned. |